Utilizing sequence intrinsic composition to classify protein-coding and long non-coding transcripts

نویسندگان

  • Liang Sun
  • Haitao Luo
  • Dechao Bu
  • Guoguang Zhao
  • Kuntao Yu
  • Changhai Zhang
  • Yuanning Liu
  • Runsheng Chen
  • Yi Zhao
چکیده

It is a challenge to classify protein-coding or non-coding transcripts, especially those re-constructed from high-throughput sequencing data of poorly annotated species. This study developed and evaluated a powerful signature tool, Coding-Non-Coding Index (CNCI), by profiling adjoining nucleotide triplets to effectively distinguish protein-coding and non-coding sequences independent of known annotations. CNCI is effective for classifying incomplete transcripts and sense-antisense pairs. The implementation of CNCI offered highly accurate classification of transcripts assembled from whole-transcriptome sequencing data in a cross-species manner, that demonstrated gene evolutionary divergence between vertebrates, and invertebrates, or between plants, and provided a long non-coding RNA catalog of orangutan. CNCI software is available at http://www.bioinfo.org/software/cnci.

برای دانلود رایگان متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

Long non-coding RNAs and their significance in human diseases

Protein-coding genes account for only a small fraction of the human genome and most of the genomic sequences are transcriptionally silent, but recent observations indicate significant functional elements, including non-coding protein transcripts in the human genome. Long non-coding RNAs (lncRNAs) have been defined as transcripts of >200 nucleotides without protein-coding capacity that perform t...

متن کامل

Phylogenetic Analysis of Three Long Non-coding RNA Genes: AK082072, AK043754 and AK082467

Now, it is clear that protein is just one of the most functional products produced by the eukaryotic genome. Indeed, a major part of the human genome is transcribed to non-coding sequences than to the coding sequence of the protein. In this study, we selected three long non-coding RNAs namely AK082072, AK043754 and AK082467 which show brain expression and local region conservation among vertebr...

متن کامل

The Roles of Long non-coding RNAs (lncRNA) in Prostate Cancer

Background & Objective: Prostate cancer is a compound condition in which gene expression has altered. Several surveys have revealed that genetic components have been involved in prostate cancer progression. Findings proposed that they can modify a noteworthy portion of disposing of elements, which is associated to the developing prostate cancer in protein coding sequences. The purpose of this r...

متن کامل

Evaluation of Long Stress-Induced Non-coding Transcripts 5 Polymorphism in Iranian Patients with Bladder Cancer

Background: Bladder cancer (BC) is the most commonly diagnosed genitourinary cancer in Iran, presented in both men and women. BC is a multifactorial trait resulting from the complex interaction between several genes and environmental factors. Long stress-induced non-coding transcript 5 (LSINCT5), a member of the long non-coding RNAs, is abundantly expressed in high proliferative cells, as well ...

متن کامل

P87: The Role of the Long Non-Coding RNA Sequences (LncRNAs) in Neurological Disorders

Precise interpretation of the transcriptome sequences in the several species showed that the major part of genome has been transcribed; however, just a few amounts of the transcription sequences have open-reading frames which are conversed during the evolution. So, it is unlikely that many of the transcribed sequences code the proteins. Among the all human non-coding transcripts, at least 10000...

متن کامل

ذخیره در منابع من


  با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

عنوان ژورنال:

دوره 41  شماره 

صفحات  -

تاریخ انتشار 2013